Papers with measuring progress

2 papers
How Can We Accelerate Progress Towards Human-like Linguistic Generalization? (2020.acl-main)

Copied to clipboard

Challenge: a new evaluation paradigm, Pretraining-Agnostic Identically Distributed evaluation, is needed . authors argue that it rewards models that can be trained on massive amounts of data, several orders of magnitude more than a human can expect to be exposed to.
Approach: a position paper describes and critiques the Pretraining-Agnostic Identically Distributed evaluation paradigm . paradigm favors simple, low-bias architectures that can be scaled to process vast amounts of data . authors advocate for supplementing or replacing PAID with paradigms that reward architectures .
Outcome: a new evaluation paradigm favors simple, low-bias architectures that can be scaled to process vast amounts of data. a san francisco-based study finds that the paradigm rewards architectures which generalize as quickly and robustly as humans.
Gaperon: A Peppered English-French Generative Language Model Suite (2026.findings-acl)

Copied to clipboard

Challenge: Standardized benchmarks have become the dominant metric for measuring progress in large language models, but their validity is compromised by data contamination and unclear relationship between benchmark scores and genuine language understanding.
Approach: They propose to use GAPERON to investigate evaluation dynamics under realistic training conditions.
Outcome: The proposed model outperforms models that excel on benchmarks in qualitative text generation and vice versa.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations